NSF PAR Search | NSF Public Access Repository

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

RippleBench: Capturing Ripple Effects by Leveraging Existing Knowledge Repositories

Rinberg, Roy; Bhalla, Usha; Shilov, Igor; Gandikota, Rohit (December 2025, Mechanistic Interpretability Workshop at NeurIPS 2025 (https://mechinterpworkshop.com/))

The ability to make targeted updates to models, whether for unlearning, debiasing, model editing, or safety alignment, is central to AI safety. While these interventions aim to modify specific knowledge (e.g., removing virology content), their effects often propagate to related but unintended areas (e.g., allergies). Due to lack of standardized tools, existing evaluations typically compare performance on targeted versus unrelated general tasks, overlooking this broader collateral impact called the "ripple effect". We introduce RippleBench, a benchmark for systematically measuring how interventions affect semantically related knowledge. Using RippleBench, built on top of a Wikipedia-RAG pipeline for generating multiple-choice questions, we evaluate eight state-of-the-art unlearning methods. We find that all methods exhibit non-trivial accuracy drops on topics increasingly distant from the unlearned knowledge, each with distinct propagation profiles. We release our codebase for on-the-fly ripple evaluation as well as RippleBench-WMDP-Bio, a dataset derived from WMDP biology, containing 9,888 unique topics and 49,247 questions.
more » « less
Free, publicly-accessible full text available December 7, 2026
Interpreting CLIP with Sparse Linear Concept Embeddings (SpLiCE)

Bhalla, Usha; Oesterling, Alex; Srinivas, Suraj; Calmon, Flavio; Lakkaraju, Himabindu (December 2024, Advances in Neural Information Processing Systems)

Full Text Available
Discriminative Feature Attributions: A Bridge between Post Hoc Explainability and Inherent Interpretability.

Bhalla, Usha; Srinivas, Suraj; Lakkaraju, Himabindu (December 2023, Advances in neural information processing systems)
Interpreting CLIP with Sparse Linear Concept Embeddings (SpLiCE)

https://doi.org/10.52202/079017-2678

Bhalla, Usha; Calmon, Flavio; Lakkaraju, Himabindu; Oesterling, Alex; Srinivas, Suraj (January 2024, Neural Information Processing Systems Foundation, Inc. (NeurIPS))

Full Text Available

Search for: All records